Papers with phrase localization
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for Grounded Multimodal Named Entity Recognition (GMNER) lack a strong correlation between image-text pairs and is ungroundable. |
| Approach: | They propose a framework that reformulates GMNER into a joint MNER-VE-VG task by leveraging large language models as a connecting bridge. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on the existing GMNER dataset and achieves absolute leads of 10.65%, 6.21%, and 8.83% in all three subtasks. |
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios. |
| Approach: | They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus. |
| Outcome: | The proposed dataset is the first multilingual image-caption dataset with Japanese translations. |